SQL injection (SQLI) remains a critical web application security threat, while conventional signature-based detection methods face limitations in identifying evolving attack patterns and handling high-dimensional web request data. This study proposes an improved machine learning framework for SQLI detection based on TF-IDF N-grams and Latent Dirichlet Allocation (LDA). HTTP requests from the CSIC-2010 and ECML/PKDD-2007 datasets are processed to extract SQLI-related traffic using SQL keyword-based filtering. TF-IDF trigram features are generated to represent request characteristics, followed by LDA-based dimensionality reduction to obtain a compact latent semantic representation. Support Vector Machine (SVM) and Random Forest (RF) classifiers are employed for classifying malicious and genuine requests. Experimental results demonstrate the effectiveness of the proposed LDA-based approach. On the CSIC-2010 dataset, the SVM and RF models achieve 99.90% and 98.13% accuracy, respectively. On the ECML/PKDD-2007 dataset, the improved framework achieves 98.70% accuracy using SVM and 99.85% using RF, with approximately 99% true-positive rate and mean ROC performance. The findings indicate that LDA-based feature reduction can improve the effectiveness and efficiency of machine-learning-based SQLI detection.
Introduction
The text presents an Machine Learning (ML)-based SQL Injection (SQLI) detection system for protecting database-driven web applications. SQL Injection is a serious cybersecurity vulnerability in which malicious input can manipulate database queries, potentially compromising confidentiality, integrity, authentication, and authorization.
Traditional SQLI detection methods generally depend on predefined signatures, rules, regular expressions, and lexical patterns. Although these methods can detect known attacks, they may struggle with new or modified SQLI patterns, require continuous signature maintenance, and can produce false alarms or application-specific limitations. Machine Learning provides an alternative by learning attack patterns from previously labeled HTTP requests and detecting malicious requests that may not exactly match predefined signatures.
Main Problem
Raw web logs contain large amounts of mixed legitimate and malicious HTTP traffic. Therefore, an effective SQLI detection system needs to:
Extract SQLI-related requests from web logs.
Balance malicious and legitimate data.
Convert textual HTTP requests into numerical features.
Reduce the resulting high-dimensional feature space.
Classify requests as SQLI or genuine.
Maintain good detection accuracy while reducing computational requirements.
Proposed Approach
The study proposes an improved framework combining:
The major improvement is replacing SVD (Singular Value Decomposition) with LDA (Latent Dirichlet Allocation) for dimensionality reduction.
TF-IDF with trigrams converts HTTP requests into numerical representations and captures important sequential patterns.
LDA reduces the large feature space by identifying latent topics and relationships between frequently co-occurring terms.
SVM (Support Vector Machine) and Random Forest (RF) are used to classify requests as malicious or legitimate.
Dataset Preparation
The research uses two publicly available datasets:
CSIC-2010
ECML/PKDD-2007
SQLI requests are extracted from the raw web logs using SQL-related keyword filtering. The malicious and genuine requests are then balanced to reduce the effect of class imbalance. For the improved ECML/PKDD-2007 experiment, the prepared dataset contains 2,274 SQLI requests and 2,274 genuine requests.
Research Gap
The study identifies several limitations in previous research:
Difficulty obtaining and preparing realistic SQLI-specific datasets.
Excessive focus on classifier selection rather than feature representation.
SVD reduces dimensionality but does not explicitly capture latent semantic relationships.
Many studies evaluate accuracy without sufficiently considering computational performance.
The proposed research addresses these issues by using LDA-based semantic dimensionality reduction after TF-IDF trigram feature extraction.
Evaluation
The models are evaluated using:
Accuracy
Precision
Recall/True Positive Rate
F1-score
ROC analysis
Training time
Testing time
The study aims to determine whether LDA can provide more informative and compact features than SVD while improving both SQLI detection performance and computational efficiency.
Conclusion
This study presented an improved machine-learning framework for detecting SQL injection attacks in HTTP web log data using TF-IDF N-gram feature representation and Latent Dirichlet Allocation (LDA). The framework addresses the high dimensionality of textual web requests by applying LDA to obtain a compact latent representation prior to classification. SVM and Random Forest were employed to evaluate the effectiveness of the resulting feature space.
Experimental results demonstrate that replacing SVD with LDA substantially improved SQLI detection performance on the ECML/PKDD-2007 dataset. The proposed approach achieved 98.70% accuracy with SVM and 99.85% with Random Forest, compared with 84.61% and 82.97%, respectively, for the earlier SVD-based configuration. The improved framework also demonstrated approximately 99% TPR and mean ROC performance, while requiring lower training and testing time.
These findings indicate that semantic-oriented dimensionality reduction can play an important role in improving ML-based SQLI detection. In particular, LDA provides a compact representation based on latent topic relationships and, for the datasets investigated, proved more effective than SVD in reducing dimensionality.
Future research can extend this framework to multiclass SQLI detection, including categories such as tautology, logically incorrect queries, union-based queries, and alternate-encoding attacks. Further investigation can also consider other web vulnerabilities, deep-learning architectures, larger labeled SQLI corpora, and integrated detection-and-prevention mechanisms.
References
[1] A. Hariyani and P. Dolia, “Comprehensive review of advanced techniques for mitigating SQL injection vulnerabilities in modern applications,” International Journal of Innovative Science and Research Technology, vol. 10, no. 3, pp. 3063–3070, Mar. 2025, doi: 10.38124/ijisrt/25mar1982.
[2] K. S. Fathi, S. I. Barakat, and A. Rezk, “An effective SQL injection detection model using LSTM for imbalanced datasets,” Computers & Security, vol. 153, Art. no. 104391, Jun. 2025, doi: 10.1016/j.cose.2025.104391.
[3] A. Hariyani and P. Dolia, “An innovative method for detecting SQLi attacks by altering SQL query attribute values,” International Journal of Advanced Computer Research, vol. 14, no. 68, pp. 89–96, Sep. 2024, doi: 10.19101/IJACR.2024.1466005.
[4] E. Casmiry, N. Mduma, and R. Sinde, “Enhanced SQL injection detection using chi-square feature selection and machine learning classifiers,” Frontiers in Big Data, vol. 8, Art. no. 1686479, Nov. 2025, doi: 10.3389/fdata.2025.1686479.
[5] R. Bak?r, “UniEmbed: A novel approach to detect XSS and SQL injection attacks leveraging multiple feature fusion with machine learning techniques,” Arabian Journal for Science and Engineering, vol. 50, no. 19, pp. 15591–15604, 2025, doi: 10.1007/s13369-024-09916-4.
[6] A. Hariyani and P. Dolia, “CryptoSQLShield: A comprehensive study on cryptography-assisted methods for SQL injection defense,” International Journal of Engineering Research & Technology (IJERT), vol. 15, no. 1, Jan. 2026, Art. no. IJERTV15IS010090, doi: 10.17577/IJERTV15IS010090.
[7] X. Wang, Y. Zheng, Z. Wan, and M. Zhang, “SVD-LLM: Truncation-aware singular value decomposition for large language model compression,” in Proc. 13th Int. Conf. Learn. Representations (ICLR), 2025.
[8] A. Damayanti and A. Baita, “Comparison of support vector machine (SVM) and random forest (RF) algorithm performance with random undersampling technique to predict gestational diabetes mellitus risk,” Journal of Applied Informatics and Computing, vol. 9, no. 2, pp. 328–337, Mar. 2025, doi: 10.30871/jaic.v9i2.9009.
[9] A. Hariyani and P. Dolia, “A cryptography-enforced SQL query integrity framework for complete SQL injection prevention,” International Journal of Scientific Research in Engineering and Management, vol. 10, no. 3, pp. 1–9, Mar. 2026, doi: 10.55041/IJSREM58457.
[10] S. Hameed, M. Nauman, N. Akhtar, M. A. B. Fayyaz, and R. Nawaz, “Explainable AI-driven depression detection from social media using natural language processing and black box machine learning models,” Frontiers in Artificial Intelligence, vol. 8, Art. no. 1627078, Sep. 2025, doi: 10.3389/frai.2025.1627078.
[11] A. Hariyani, “CDiCENet: A structure-aware lightweight deep learning model for SQL injection attack detection on resource-constrained devices,” International Journal of Engineering Research & Technology (IJERT), vol. 15, no. 8, Aug. 2026, Art. no. IJERTV15IS080682, doi: 10.17577/IJERTV15IS080682.
[12] M. E. D. Rafi, M. H. M. Fajar, M. S. Purwanto, A. Hilyah, A. S. Bahri, and H. K. Rahayu, “Analysis of formation Ronggojalu spring and Probolinggo active fault continuity with satellite data gravity method,” Jurnal Penelitian Pendidikan IPA, vol. 9, no. 10, pp. 8456–8461, Oct. 2023, doi: 10.29303/jppipa.v9i10.3399.
[13] A. Hariyani, “A structured framework for detection and mitigation of SQL injection vulnerabilities in web applications,” International Journal of Creative Research Thoughts (IJCRT), vol. 14, no. 8, pp. e332–e343, Aug. 2026, Paper ID: IJCRT2608464.
[14] A. Hariyani and P. Dolia, “SecureSQL: Preventing SQL injection attacks using cryptographic query protection and intelligent detection,” in Proc. 4th Int. Conf. Cybersecurity and Generative Artificial Intelligence (CyberGenAI’2026), Mar. 2026, doi: 10.5281/zenodo.18957529.
[15] P. Guleria, J. Frnda, and P. N. Srinivasu, “NLP based text classification using TF-IDF enabled fine-tuned long short-term memory: An empirical analysis,” Array, vol. 27, Art. no. 100467, Sep. 2025, doi: 10.1016/j.array.2025.100467.
[16] S. He, Y. Zhang, D. Liang, and P. K. Sharma, “An unsupervised malicious web request detection based on transformer and contrastive learning,” IEEE Transactions on Network and Service Management, vol. 22, no. 4, pp. 3281–3294, Aug. 2025, doi: 10.1109/TNSM.2025.3563089.